Papers with composing evaluation guidelines

1 papers
Defining and Detecting Vulnerability in Human Evaluation Guidelines: A Preliminary Study Towards Reliable NLG Evaluation (2024.naacl-long)

Copied to clipboard

Challenge: Unreliable evaluation guidelines can yield inaccurate assessment outcomes, potentially impeding the advancement of NLG in the right direction.
Approach: They propose to collect annotated human evaluation guidelines and a method for detecting guideline vulnerabilities using Large Language Models.
Outcome: The proposed dataset includes eight vulnerabilities and a method for detecting guideline vulnerabilities.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations